Papers with reference-free metrics

9 papers
Fusion-Eval: Integrating Assistant Evaluators with LLMs (2024.emnlp-industry)

Copied to clipboard

Challenge: Recent studies have employed large language models (LLMs) as reference-free metrics for NLG evaluation, enhancing adaptability to new tasks tasks.
Approach: They propose a method that leverages large language models to integrate insights from various assistant evaluators.
Outcome: The proposed approach achieves a 0.962 system-level Kendall-Tau correlation with humans on SummEval and a 0.7444 turn-level Spearman correlation on TopicalChat, which is significantly higher than baseline methods.
G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment (2023.emnlp-main)

Copied to clipboard

Challenge: Conventional reference-based metrics have low correlation with human judgments, especially for open-ended generation tasks.
Approach: They propose to use large language models as reference-free NLG evaluators to assess the quality of NLG outputs.
Outcome: The proposed framework outperforms all previous methods in two generation tasks, and has a Spearman correlation of 0.514 with human on summarization task, and a large variance in human judgments.
CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text Generation (2022.acl-long)

Copied to clipboard

Challenge: Existing reference-free metrics have obvious limitations for evaluating controlled text generation models.
Approach: They propose an unsupervised reference-free metric which evaluates controlled text generation from different aspects by formulating each aspect into multiple text infilling tasks.
Outcome: The proposed metric has higher correlations with human judgments while obtaining better generalization of evaluating generated texts from different models and with different qualities.
On the Evaluation Metrics for Paraphrase Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for paraphrase generation are not designed for the task, but adopted from other evaluation tasks.
Approach: They propose a new evaluation metric for paraphrase generation that uses reference-based and reference-free metrics.
Outcome: The proposed evaluation metric outperforms existing metrics and is more reliable than reference-based metrics.
Is Reference Necessary in the Evaluation of NLG Systems? When and Where? (2024.naacl-long)

Copied to clipboard

Challenge: Despite recent advances in reference-free metrics, it has not been well understood when and where they can be used as an alternative to reference-based metrics.
Approach: They propose to use reference-free metrics to evaluate NLG systems . they find they have a higher correlation with human judgment and greater sensitivity to deficiencies in language quality .
Outcome: The proposed metrics exhibit higher correlation with human judgment and greater sensitivity to deficiencies in language quality.
A Quality-based Syntactic Template Retriever for Syntactically-Controlled Paraphrase Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing syntactically-controlled paraphrase generation models perform well with human-annotated or well-chosen syntaktic templates.
Approach: They propose a quality-based Syntactic Template Retriever to retrieve templates based on the quality of the to-be-generated paraphrases.
Outcome: The proposed algorithm can generate high-quality paraphrases without sacrificing quality.
AdParaphrase v2.0: Generating Attractive Ad Texts Using a Preference-Annotated Paraphrase Dataset (2025.findings-acl)

Copied to clipboard

Challenge: Identifying factors that make ad text attractive is essential for advertising success . identifying the linguistic factors presents a significant challenge because of the intricate interplay between the semantic content and its linguistic expression.
Approach: They propose to use a dataset for ad text paraphrasing that contains human preference data to enable analysis of linguistic factors.
Outcome: The proposed dataset is 20 times larger than v1.0 and contains 16,460 pairs of ad text paraphrase pairs . it shows that human preference and ade- t attractiveness are related .
Reference-Free Evaluation of Taxonomies (2026.findings-acl)

Copied to clipboard

Challenge: Taxonomies are used to classify items, ideas or organisms based on shared characteristics.
Approach: They introduce two reference-free metrics for quality evaluation of taxonomies in the absence of labels.
Outcome: The proposed metrics correlate well with F1 against ground truth taxonomies on five taxonomies and improve hierarchical classification when used with label hierarchies.
When Cohesion Lies in the Embedding Space: Embedding-Based Reference-Free Metrics for Topic Segmentation (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in topic segmentation have led to a surge in interest in reference-free metrics, designed to score a hypothesised segmentation of a document without the need to refer to any expert annotation.
Approach: They propose a common framework for reference-free topic segmentation metrics and a new method for the embedding space.
Outcome: The proposed framework outperforms existing metrics based on human annotations while allowing for conversational data to outperformed other metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations